Blind normalization of speech from different channels and speakers
نویسنده
چکیده
This paper describes representations of time-dependent signals that are invariant under any invertible time-independent transformation of the signal time series. Such a representation is created by rescaling the signal in a non-linear dynamic manner that is determined by recently encountered signal levels. This technique may make it possible to normalize signals that are related by channel-dependent and speaker-dependent transformations, without having to characterize the form of the signal transformations, which remain unknown. The technique is illustrated by applying it to the time-dependent spectra of speech that has been filtered to simulate the effects of different channels. The experimental results show that the rescaled speech representations are largely normalized (i.e., channelindependent), despite the channel-dependence of the raw (unrescaled) speech.
منابع مشابه
تخمین سریع ضرایب پیچش در هنجارسازی طول مجرای صوتی با استفاده از امتیاز به دست آمده از مدلسازی تشخیص جنسیت
The performance of automatic speech recognition (ASR) systems is adversely affected by the variations in speakers, audio channels and environmental conditions. Making these systems robust to these variations is still a big challenge. One of the main sources of variations in the speakers is the differences between their Vocal Tract Length (VTL). Vocal Tract Length Normalization (VTLN) is an effe...
متن کاملA Pragmatic Study of Requestive Speech Act by Iranian EFL Learners and Canadian Native Speakers in Hotels
This study was an attempt to shed light on the use of requestive speech act by Iranian nonnative speakers (NNSs) of English and Canadian native speakers (NSs) of English to find out the (possible) similarities and/or differences between the request realizations, and to investigate the influence of the situational variables of power, distance, context familiarity, and L1’s (possible) influence. ...
متن کاملEvolutive Speaker Segmentation using a Repository System
When performing blind speaker segmentation one of the main problems is not knowing how many speakers appear in a conversation and wether they appear once or more than once. In this paper, an iterative method, which is based on the EvolutiveHMM is presented. Two main improvements to this system are introduced. On one hand, a repository generic speaker is used to model all utterances and all spea...
متن کاملEvolutive speaker segmentation using a repository system
When performing blind speaker segmentation one of the main problems is not knowing how many speakers appear in a conversation and wether they appear once or more than once. In this paper, an iterative method, which is based on the EvolutiveHMM is presented. Two main improvements to this system are introduced. On one hand, a repository generic speaker is used to model all utterances and all spea...
متن کاملExploring Pragmalinguistic and Sociopragmatic Variability in Speech Act Production of L2 Learners and Native Speakers
The pragmalinguistic and sociopragmatic aspects of language use vary across different situations, languages, and cultures. The separation of these two facets of language use can help to map out the socio-cultural norms and conventions as well as the linguistic forms and strategies that underlie the pragmatic performance of different language speakers in a variety of target language use situatio...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
- CoRR
دوره cs.CL/0204003 شماره
صفحات -
تاریخ انتشار 2002